{"id":"277bab36-ecae-4e49-a4ec-04d8196c52d9","arxiv_id":"2608.04646","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.","lead":"This paper tests whether newer 'reasoning' AI models are better at understanding what other people believe, and whether their apparent gains come from a new ability or from being more consistent under changes in wording. The authors report that reasoning models are more robust to prompt and task variations, and interpret this as support for a robustness-based account rather than a new Theory of Mind capability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The robustness conclusion rests on one uncontrolled vendor comparison (Claude thinking on/off), and the paper's own Table 6 shows item-level counterexamples, so the central claim is not yet distinguishable from API or general-capability confounds.","rationale":"The paper is a carefully hedged behavioral study. Its strongest claim is that reasoning models show increased robustness to prompt and task variation, and that this supports a robustness-based account rather than a new ToM-specific ability. For that claim to hold, the observed differences must be attributable to the reasoning process rather than to confounds such as API filtering, model version, temperature, or interface artifacts. The reader's weakest assumption identifies exactly this condition, and the paper's own Section 3.1 and Section 5.2 concede that the comparison is not fully controlled. I agree with that assessment. My independent check of Table 6 reinforces the concern: the item-level pattern is not uniformly in favor of reasoning models, and the paper reports no sample sizes or error bars, so 'consistently' overstates the evidence. The proposed concrete test—replicating the core comparison with an open-weight reasoning model and its accessible base model under matched decoding—would settle whether the Claude thinking-on/off gap is a genuine effect of reasoning or an artifact of the vendor configuration. I give credit for the paper's transparency, its explicit limitations, and its qualitative analysis of reasoning traces, all of which are useful despite the underdetermination. The result is a plausible hypothesis, not a demonstrated conclusion, so the reader's CONDITIONAL verdict remains appropriate without requiring a change.","tokens_in":12279,"tokens_out":7881,"duration_ms":86748,"concrete_test":"Run the identical 14 modified tasks from Table 6 on DeepSeek-R1 and its accessible base model DeepSeek-V3, using the same library, identical prompts, and matched decoding settings, with temperature 1 and at least 20 independent samples per item. If R1 does not show the same robustness advantage over V3 that Claude thinking-on shows over Claude thinking-off—or if the advantage disappears under matched decoding—the central attribution to reasoning is not supported and the conclusion should be downgraded to an untested hypothesis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that reasoning models 'consistently exhibit increased robustness'—depends on attributing the Table 6 gap between Claude with thinking on and thinking off to the reasoning process. That attribution is not secured. Section 3.1 concedes that 'comparisons are not perfectly controlled due to API constraints; results should be interpreted as behavioral, not architectural,' and Section 5.2 concedes that Claude with thinking off 'is not the same thing as its base model.' Thinking-off Claude is a different product configuration that may differ in system prompt, safety filtering, response formatting, and post-processing; it is not the same model with inference-time scaling removed. The only within-family comparison is therefore one uncontrolled pair. Moreover, the paper provides no sample sizes, error bars, or statistical test for Table 6, and the item-level pattern is not 'consistent': on 2A and 2B all models score 0; GPT-5-high (reasoning) scores 0 on 1B.2 where non-reasoning Claude scores 0.5; R1 scores 0.5 on 1C.1 where non-reasoning Claude scores 1.0. Higher average accuracy on modified items is compatible with better instruction following or higher general capability, not uniquely with a robustness-based account. The qualitative traces are suggestive but cannot rule out meta-knowledge or interface artifacts, as the paper itself notes. Thus the load-bearing assumption is that the observed gap is caused by reasoning; without a controlled replication, the central interpretation remains underdetermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates recent reasoning-oriented LLMs (GPT-5, Claude, R1, Grok-3-mini) on Theory of Mind (ToM) tasks from the battery of van Duijn et al., together with prompt perturbations inspired by Ullman, and compares reasoning-enabled configurations against non-reasoning configurations where available. It reports near-ceiling performance on Sally-Anne, Strange Stories, and Imposing Memories tasks, and mixed but generally higher scores for reasoning models on a set of modified simple ToM prompts, with a thinking-off Claude condition scoring lower on several items. The authors interpret the overall pattern as evidence that reasoning models exhibit increased robustness to prompt and task variation, supporting a robustness-based account of recent ToM gains rather than a new ToM-specific ability. The paper explicitly disclaims causal attribution to RLVR training and describes the comparisons as behavioral rather than fully controlled.","tokens_in":12572,"tokens_out":3334,"duration_ms":38722,"significance":"If the central interpretation holds, the paper makes a timely contribution to the debate on whether recent ToM improvements in LLMs reflect genuinely new social-cognitive capabilities or improved stability in reaching solutions that were already within reach. The manuscript is transparent: it provides open code, data, and qualitative reasoning traces; it incorporates external benchmark results from prior work; and it repeatedly hedges its causal claims. For these reasons, the paper is a useful contribution even though the evidence remains suggestive rather than conclusive. The significance would be strengthened if the quantitative support for the main claim matched the level of the abstract's assertion.","major_comments":[{"comment":"This is the load-bearing comparison for the paper's interpretation. The current text openly acknowledges the confound, but then proceeds to draw the robustness conclusion from the very same comparison. A reader cannot verify that the observed differences are due to reasoning rather than to uncontrolled configuration differences. The authors should either run a controlled study (e.g., same base model with and without inference-time scaling, if accessible) or reframe the conclusion as a hypothesis supported only by suggestive evidence, not as an observed regularity.","section":"Section 3.1 and Table 6"},{"comment":"The absence of inferential statistics is especially problematic because the differences in Table 6 are small and the item counts appear to be low (e.g., 2C.1 and 2C.2 appear twice, suggesting a small item set). The paper's own phrasing in Section 5.2 acknowledges that 'no clear quantification of such an effect can be provided,' which is at odds with the abstract's unqualified claim of consistent increased robustness. This mismatch should be resolved.","section":"Section 4 and Tables 3-6"},{"comment":"The paper's own limitations section is admirably candid, but the abstract and conclusion are not fully aligned with those caveats. Since the central theoretical contribution is the robustness-based interpretation, the mismatch between the hedged body and the assertive abstract is a load-bearing presentation issue that should be corrected.","section":"Section 5.1 and Abstract"}],"minor_comments":[{"comment":"The column layout of Table 1 is difficult to read because the benchmark names and the three sub-columns (ParaphrasedToMi, FANToM, MMToM-QA) are not separated clearly; the numbers for the different benchmarks run together. Reformatting the table with explicit sub-headers or separate rows would improve readability.","section":"Table 1"},{"comment":"The column header 'whitelie' should be 'white lie' for consistency with the prose.","section":"Table 4"},{"comment":"The phrase 'significantly better' in the second paragraph of Section 4.4 should be replaced with a descriptive term such as 'numerically higher,' given that no statistical test is reported.","section":"Section 4.4"},{"comment":"The qualitative analysis of reasoning traces is not systematic: the criteria for identifying 'perspective-taking steps' or 'meta-knowledge' are not defined, and it is unclear how many responses were inspected and by whom. A brief coding scheme or inter-annotator agreement would improve the credibility of these qualitative claims.","section":"Section 4.5"},{"comment":"The citation for the 'orders of reasoning' concept [12] is appropriate, but the definition in Section 2.2 is informal; a more precise reference to the recursive structure, or a page/section number in [12], would help readers.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and well-hedged, but the abstract overstates the robustness finding relative to the evidence. The core issue is that the only within-family comparison (Claude thinking on/off) is uncontrolled, and no inferential statistics support the central claim. I believe the paper can be made acceptable by adding item-level counts and confidence intervals, or by toning down the abstract and conclusion to explicitly say the evidence is 'consistent with, but not decisive for, a robustness-based account.' If the authors choose the latter route, the contribution becomes a preliminary behavioral study rather than a definitive finding, which may affect its fit for the journal. I would lean toward accepting after a revision that brings the claims in line with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first direct behavioral comparison of reasoning-enabled and non-reasoning models on established psychological ToM tests, with new task-preserving prompt perturbations. The robustness-over-capability framing is borrowed from Yue et al., but the application to ToM is new, and the authors are refreshingly transparent about what they can and cannot claim. The central conclusion—that gains reflect robustness rather than new ToM-specific ability—is plausible but not actually pinned down by the data, and the paper's own caveats say as much.\n\nWhat is genuinely good: the authors run a standard psychological battery (Sally-Anne, Strange Stories, Imposing Memories, plus modifications of Ullman-style prompts) on current reasoning models, and they compare thinking-on and thinking-off Claude where possible. They release prompts, data, and code, and they score outputs with a fixed rubric inherited from prior work. They also bring in external benchmark numbers from Kim et al. rather than relying only on their own runs. The qualitative trace analysis is a plus: they actually look at model reasoning, identify distinct failure modes, and note when meta-knowledge of the task might be doing the work. The discussion section is honest about the limits: no causal identification of RLVR, no access to base models, and the possibility of spurious correlations.\n\nThe soft spots are real but not hidden. The stress-test note is on target: the load-bearing comparison is Claude with thinking on versus thinking off, and that is an uncontrolled product difference, not a clean manipulation of reasoning. Section 5.2 concedes the thinking-off model is not the base model. Table 6 also does not show the consistent pattern the abstract claims: all models score 0 on the 2A/2B tasks, and there are item-level inversions (e.g., R1 at 0.5 on 1C.1 while no-reas-claude gets 1.0). There are no error bars or statistical tests anywhere, so we cannot tell how much of the gap is noise. The authors deserve credit for using hedged language, but the word 'consistently' in the abstract overstates what the table shows. These are addressable: report sample sizes and confidence intervals, soften the abstract, and ideally add a comparison with a real base model (e.g., R1 vs V3-base) or a controlled API configuration.\n\nBottom line: this is a useful, honest empirical contribution that will matter to anyone designing ToM evaluations. The central claim should be read as a well-articulated hypothesis rather than a proven result. I would cite it for the benchmark comparison, and I would bring it to a reading group for the methodological discussion. A serious referee should see it; with revisions addressing the statistical reporting and the abstract's overreach, it would be a solid paper.","headline":"Useful, honest ToM evaluation with a plausible robustness story, but the central comparison is an uncontrolled vendor config and the abstract overstates the evidence.","tokens_in":13076,"tokens_out":2838,"would_cite":true,"duration_ms":29647,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reasoning-enabled language models are more robust on Theory of Mind tasks, and the paper argues this robustness, not a new mental-state ability, explains recent gains.","keywords":["Theory of Mind","reasoning models","reinforcement learning with verifiable rewards","prompt robustness","false-belief tests","LLM evaluation","chain-of-thought reasoning","benchmark perturbation"],"falsifier":"Re-run the modification battery on Claude with thinking on and off across a much larger set of task-preserving paraphrases and multiple seeds; the robustness claim predicts lower per-paraphrase accuracy variance and higher minimum accuracy for thinking-on, so matching or better performance from the thinking-off configuration would contradict it.","tokens_in":12068,"feed_emoji":"🧠","tokens_out":11149,"duration_ms":113039,"temperature":0.7,"pith_summary":"This paper asks whether the strong recent performance of reasoning-oriented large language models on Theory of Mind (ToM) tests marks a new social-cognitive ability or a narrower improvement. Its answer is that the gains are best read as robustness: models trained to spend inference time on a chain of thought reach a correct answer that the underlying model could already, in principle, reach, and they do so more reliably when the same task is rephrased or perturbed. The evidence comes from adapted psychological tests (first- and second-order false belief, Strange Stories, Imposing Memories), prompt variants designed to disturb reasoning, a thinking-on versus thinking-off comparison, and reanalysis of third-party benchmarks. If the paper is right, progress on ToM benchmarks should be tracked by stability under variation, not just by average accuracy.","feed_headline":"Mind-reading scores rise via robustness, not new ability","feed_subtitle":"Prompt-variation tests show that stability, not a new ability, explains recent mind-reading gains.","key_machinery":"The machinery is paired variation and consistency scoring. Task-preserving prompt modifications, inspired by earlier demonstrations that GPT-3 failed trivial rephrasings, change surface form while leaving the false-belief content intact; comparing Claude with thinking on and off on these variants isolates the effect of the reasoning process as far as the interface allows; and FANToM's all-questions-correct scoring turns answer instability into a hard penalty, making robustness differences visible in benchmark numbers. The paper also scores answers and their reasoning together on a 0-1-2 scale, so a correct inference path counts as evidence of stability rather than only a correct label.","core_discovery":"Reasoning-oriented LLMs consistently show increased robustness to prompt variations and task perturbations on ToM material. On the authors' tests, thinking-enabled Claude outperformed its thinking-off counterpart and was markedly better on the modified prompts designed to derail reasoning; GPT-5 made only a single partial mistake across the full battery; and third-party benchmark results show reasoning models above their non-reasoning peers, with the largest gap on FANToM, a benchmark that only credits a model when it answers every question of a given type correctly. The authors read this pattern as support for a robustness-based account: reasoning training stabilizes the selection of an already-available inference path rather than adding a fundamentally new capacity to represent mental states. They explicitly restrict the claim to behavior, noting that the comparisons are not perfectly controlled and do not causally isolate reinforcement-learning training.","pith_inferences":["A natural extension is to fold robustness into the operational definition of a model skill: if a capability disappears under task-preserving variation, it is arguably not a stable capability, making robustness a component of ability rather than a separate evaluation axis.","The account predicts an inverse relationship between reasoning effort and answer variance across paraphrases; this can be tested by sampling many paraphrases at increasing reasoning budgets and measuring variance in the final answer.","The same robustness lens could apply outside ToM: RLVR models' gains on math and code benchmarks may likewise reflect a narrowed, more consistent solution distribution rather than new knowledge, and paraphrased versions of those problems would expose this.","A decisive version of the comparison would pair each reasoning model with its exact base model on the same perturbation battery; the authors note this is feasible for models like R1 versus its base, and doing so would turn the behavioral pattern into a cleaner causal test."],"forward_implications":["ToM evaluation should report accuracy across task-preserving prompt variants, since average benchmark scores alone cannot separate robustness from newly added capability.","Earlier prompt-sensitivity failures, such as GPT-3's collapse under trivial alterations, can be reinterpreted as failures to reliably reach an inference path the model already had, rather than as proof that ToM competence was absent.","Benchmarks that penalize inconsistency, like all-questions-correct scoring, are more sensitive detectors of the reasoning models' advantage than benchmarks that average over independent questions.","If the account holds, reasoning-oriented models should be expected to behave more consistently in social and agentic settings under rephrased inputs, while their ceiling remains bounded by the base model."],"supporting_citations":[{"why":"Supplies the psychological test battery (Sally-Anne, Strange Stories, Imposing Memories) and the 2023 baseline against which the new models are measured.","marker":"[6]"},{"why":"Supplies the trivial prompt alterations that GPT-3 failed, motivating the modification tasks and the robustness interpretation.","marker":"[24]"},{"why":"Supplies third-party ToM benchmark results for reasoning and non-reasoning models, including FANToM, whose consistency penalty shows the largest gap.","marker":"[9]"},{"why":"Supplies the finding that reinforcement-learned reasoning remains bounded by the base model, grounding the reading that reasoning stabilizes existing capabilities.","marker":"[28]"},{"why":"Introduces FANToM, the benchmark whose all-questions-correct scoring penalizes inconsistency and thus magnifies the reasoning advantage.","marker":"[10]"}],"fun_headline_variants":["ToM gains trace to robustness, not new ability","Reasoning models: stronger ToM via robustness","Robustness explains AI mind-reading gains","Stable answers drive ToM gains, not new insight"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's interpretation assumes that the differences between thinking-on and thinking-off Claude, and between reasoning and non-reasoning models more broadly, come from the reasoning process itself rather than from API filtering, model version, temperature, or other uncontrolled interface factors; the authors concede the comparisons are behavioral, not perfectly controlled.","fun_headline_variants_meta":{"raw":{"variants":["ToM gains trace to robustness, not new ability","Reasoning models: stronger ToM via robustness","Robustness explains AI mind-reading gains","Stable answers drive ToM gains, not new insight"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00014,"raw_usage":{"total_tokens":1104,"prompt_tokens":830,"completion_tokens":274,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":213}},"tokens_in":446,"tokens_out":274,"duration_ms":3490,"temperature":1.0,"reasoning_tokens":213,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:52:47.116510+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the modification battery on Claude with thinking on and off across a much larger set of task-preserving paraphrases and multiple seeds; the robustness claim predicts lower per-paraphrase accuracy variance and higher minimum accuracy for thinking-on, so matching or better performance from the thinking-off configuration would contradict it.","supporting_citations":[{"cited_title":"In: COLM (2025)","cited_arxiv_id":null,"evidence_quote":"Supplies third-party ToM benchmark results for reasoning and non-reasoning models, including FANToM, whose consistency penalty shows the largest gap."}],"review_version":1}